Papers with visual processing
“I’ve Seen Things You People Wouldn’t Believe”: Hallucinating Entities in GuessWhat?! (2021.acl-srw)
Copied to clipboard
| Challenge: | a problem with natural language generation systems is the generation of tokens that are unrelated to the source input. |
| Approach: | They propose two new models to play the GuessWhat?! referential game . they propose to adapt the best visual processing models available to mitigate this issue . |
| Outcome: | The proposed models generate few hallucinations compared to other models available in the literature. |
Picturing Ambiguity: A Visual Twist on the Winograd Schema Challenge (2024.acl-long)
Copied to clipboard
| Challenge: | Large Language Models have demonstrated remarkable success in tasks like the Winograd Schema Challenge (WSC), showcasing advanced textual common-sense reasoning. |
| Approach: | They propose a framework to isolate models' ability in pronoun disambiguation from other visual processing challenges. |
| Outcome: | The proposed framework isolates the models’ ability in pronoun disambiguation from other visual processing challenges. |
GlyphPattern: An Abstract Pattern Recognition for Vision-Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for abstract pattern recognition are easier because they do not involve a natural language description of the pattern. |
| Approach: | They present a dataset that pairs human-written descriptions of visual patterns with three visual presentation styles. |
| Outcome: | The proposed benchmark pairs human-written and human-verified patterns with three visual presentation styles. |
Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human Gaze (2020.emnlp-main)
Copied to clipboard
| Challenge: | a long tradition of cognitive studies shows that the interplay between language and vision is complex. |
| Approach: | They propose an approach to image description generation where visual processing is modelled sequentially. |
| Outcome: | The proposed model exploits gaze-driven attention to produce better descriptions . it sheds light on human cognitive processes by comparing different ways of aligning gaze with language production. |
LLMs Can Compensate for Deficiencies in Visual Representations (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a strong language backbone in vision-language models compensates for weak visual features by contextualizing or enriching them. |
| Approach: | They investigate whether strong language backbone compensates for weak visual features . they use CLIP-based vision encoders to perform controlled self-attention ablations . |
| Outcome: | The proposed model compensates for weak visual features by contextualizing or enriching them. |